Papers with corpus annotation
A Simple yet Efficient Prompt Compression Method for Text Classification Data Annotation Using LLM (2025.coling-industry)
Copied to clipboard
| Challenge: | Existing methods to improve the accuracy of large language models (LLMs) are often impractical due to high costs and time consumption. |
| Approach: | They propose a method that uses keyword extraction to reduce prompt tokens in text annotation tasks. |
| Outcome: | The proposed method reduces prompt tokens while maintaining high accuracy. |
The ISO Standard for Dialogue Act Annotation, Second Edition (2020.lrec-1)
Copied to clipboard
Harry Bunt, Volha Petukhova, Emer Gilmartin, Catherine Pelachaud, Alex Fang, Simon Keizer, Laurent Prévot
| Challenge: | ISO standard 24617-2 for dialogue act annotation has been used in corpus annotation and in the design of components for spoken and multimodal interactive systems. |
| Approach: | ISO standard 24617-2 for dialogue act annotation is proposed for a second edition . this second edition allows a more accurate annotation of dependence relations and rhetorical relations in dialogue. |
| Outcome: | The proposed second edition of ISO 24617-2 for dialogue act annotation addresses some inaccuracies and undesirable limitations. |
Moving TIGER beyond Sentence-Level (L18-1)
Copied to clipboard
| Challenge: | TIGER 2.2-doc is a new set of annotations for the German TIger corpus. |
| Approach: | They propose a new set of annotations for the German TIGER corpus . they introduce new document-level annotations: authors and their gender. |
| Outcome: | The new annotations improve the TIGER corpus and its structure and authors and gender. |
Arabic Speech Rhythm Corpus: Read and Spontaneous Speaking Styles (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Arabic speech recordings has been built to allow comparisons between Arabic and other languages. |
| Approach: | They propose to build a corpus of Arabic speech recordings that can be compared with other languages. |
| Outcome: | The proposed corpus can be used for forensic phonetic research and casework applications. |
Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville) (L18-1)
Copied to clipboard
Annie Rialland, Martine Adda-Decker, Guy-Noël Kouarata, Gilles Adda, Laurent Besacier, Lori Lamel, Elodie Gauthier, Pierre Godard, Jamison Cooper-Leavitt
| Challenge: | BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages. |
| Approach: | This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project. |
| Outcome: | The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants. |
Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser Development (2024.lrec-main)
Copied to clipboard
Maciej Ogrodniczuk, Aleksandra Tomaszewska, Daniel Ziembicki, Sebastian Żurowski, Ryszard Tuora, Aleksandra Zwierzchowska
| Challenge: | The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation. |
| Approach: | They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework. |
| Outcome: | The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators. |